The Bivariate 2-Poisson Model for IR
نویسندگان
چکیده
Harter’s 2-Poisson model of Information Retrieval is a univariate model of the raw term frequencies, that does not condition the probabilities on document length [2]. A bivariate stochastic model is thus introduced to extend Harter’s 2-Poisson model, by conditioning the term frequencies of the document to the document length. We assume Harter’s hypothesis: the higher the probability f(X = x|L = l) of the term frequency X = x is in a document of length l, the more relevant that document is. The new generalization of the 2-Poisson model has 5 parameters that are learned term by term through the EM algorithm over term frequencies data. We explore the following frameworks:
منابع مشابه
Estimation of Count Data using Bivariate Negative Binomial Regression Models
Abstract Negative binomial regression model (NBR) is a popular approach for modeling overdispersed count data with covariates. Several parameterizations have been performed for NBR, and the two well-known models, negative binomial-1 regression model (NBR-1) and negative binomial-2 regression model (NBR-2), have been applied. Another parameterization of NBR is negative binomial-P regression mode...
متن کاملSeasonal Influences on Different Stages of In Vitro Fertilization: Stimulation and Fertilization
Background Effect of seasonal changes in human reproduction has been intensively researched. Some studies acknowledge influences of seasonal variation on natural conception, while others can not confirm them. The aim of this study was to investigate the effect of seasonal changes on different stages of assisted reproductive technology (ART) in Tehran. MaterialsAndMethods This study was carried ...
متن کاملOn Bivariate Generalized Exponential-Power Series Class of Distributions
In this paper, we introduce a new class of bivariate distributions by compounding the bivariate generalized exponential and power-series distributions. This new class contains the bivariate generalized exponential-Poisson, bivariate generalized exponential-logarithmic, bivariate generalized exponential-binomial and bivariate generalized exponential-negative binomial distributions as specia...
متن کاملModeling animal-vehicle collisions using diagonal inflated bivariate Poisson regression.
Two types of animal-vehicle collision (AVC) data are commonly adopted for AVC-related risk analysis research: reported AVC data and carcass removal data. One issue with these two data sets is that they were found to have significant discrepancies by previous studies. In order to model these two types of data together and provide a better understanding of highway AVCs, this study adopts a diagon...
متن کاملطراحی شبکه عصبی مصنوعی برای پیشبینی توأم سندرم متابولیک و شاخص مقاومت به انسولین (HOMA-IR): مطالعه قند و لیپید تهران
Background & Objective: Mixed outcomes arise when, in a multivariate model, response variables measured on different scales such as binary and continuous. In a bivariate modeling, when there are mixed response variables, the common methods in classic statistics have shortcomings. This study aimed at designing an appropriate ANN model for modeling and predicting the bivariate mixed responses i...
متن کاملذخیره در منابع من
با ذخیره ی این منبع در منابع من، دسترسی به آن را برای استفاده های بعدی آسان تر کنید
عنوان ژورنال:
دوره شماره
صفحات -
تاریخ انتشار 2013